Papers with Vision Language Models

34 papers
Beyond Visual Understanding Introducing PARROT-360V for Vision Language Model Benchmarking (2025.coling-industry)

Copied to clipboard

Challenge: Current benchmarks for evaluating Vision Language Models (VLMs) often fail to thoroughly assess these models’ abilities to understand complex visual and textual content.
Approach: They propose a benchmark that features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks.
Outcome: The PARROT-360V Benchmark features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks.
CoLLaVO: Crayon Large Language and Vision mOdel (2024.findings-acl)

Copied to clipboard

Challenge: Existing Large Language Models (LLMs) and instruction tuning have been used to drive the evolution of Vision Language Model (VLM) towards a versatile general-purpose model.
Approach: They propose a learning strategy of Dual QLoRA to preserve object-level image understanding without forgetting it during visual instruction tuning, thereby achieving a significant leap in numerous VL benchmarks in a zero-shot setting.
Outcome: The proposed model outperforms closed-source models on vision language tasks and achieves a significant leap in numerous benchmarks.
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine (2025.naacl-long)

Copied to clipboard

Challenge: Vision Language Models (VLMs) have demonstrated promise in generating visually grounded responses, but their application in the medical domain is hindered by unique challenges.
Approach: They propose a vision language model with versatile visual grounding for medicine that generates semantic segmentation masks and instance-level bounding boxes.
Outcome: The proposed model can generate semantic segmentation masks and instance-level bounding boxes, and accommodates various imaging modalities, including both 2D and 3D data.
MATE: Meet At The Embedding - Connecting Images with Long Texts (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Vision Language Models (VLMs) focus on aligning images with short descriptive captions.
Approach: They propose a method that combines VLMs with Large Language Models to efficiently align images with long texts without additional text pairs.
Outcome: The proposed method bridges the gap between VLM and LLM without additional image-long text pairs.
CAST: Cross-modal Alignment Similarity Test for Vision Language Models (2025.coling-main)

Copied to clipboard

Challenge: Vision Language Models (VLMs) are typically evaluated with Visual Question Answering tasks which assess a model’s understanding of scenes.
Approach: They propose to use visual question answering (VQA) to assess a model's understanding of scenes to probe for self-consistency across modalities.
Outcome: The proposed test does not focus on objective accuracy but rather on whether VLMs are internally consistent in their outputs.
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent Vision Language Models (VLMs) have shown tremendous promise in a wide range of realworld applications, but their size has made at-scale deployment and operation challenging due to high consumption of cloud computing resource, high latency, and expensive API calls.
Approach: They propose a master–apprentice framework for collaborative inference between large and small vision language models.
Outcome: The proposed framework improves reasoning performance on widely-recognized and challenging general reasoning benchmarks and specifically boosts reasoning of apprentice VLMs by 36.6%.
Ask Me Again Differently: GRAS for Measuring Bias in Vision Language Models on Gender, Race, Age, and Skin Tone (2026.findings-eacl)

Copied to clipboard

Challenge: Using vision language models, we examine demographic biases in VLMs across gender, race, age, and skin tone.
Approach: They propose a benchmark for uncovering demographic biases in Vision Language Models . they propose 'Gras Bias Score' to quantify bias in VLMs based on gender, race, age and skin tone .
Outcome: The proposed model achieves 98, far from the unbiased ideal of 0.
Defeating Cerberus: Privacy-Leakage Mitigation in Vision Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Existing models that process multiple modalities of data have been used for multimodal tasks, but their advanced capabilities raise privacy concerns.
Approach: They propose a method to modify the model’s internal states associated with PII-related content and to reduce the risk of PI I leakage by modifying the model's internal state.
Outcome: The proposed method achieves on average 93.3% refusal rate for various PII-related tasks with minimal impact on unrelated model performances.
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models and Vision Language Model (VLMs) have demonstrated aptitude as potential substitutes for human participants in psycholinguistic experiments.
Approach: They examine whether large language models and vision language models implicitly understand sound-based phenomena via orthography and imagery alone.
Outcome: The proposed models demonstrate sound symbolism and ability to "hear" using language and vision modules.
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines (2025.naacl-long)

Copied to clipboard

Challenge: Vision Language Models struggle with cultural-specific knowledge, especially in languages other than English and in underrepresented cultural contexts.
Approach: They propose a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects and a training dataset.
Outcome: The proposed model performs better with correct location context, but struggles with adversarial contexts and predicting specific regional cuisines and languages.
Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping (2025.naacl-long)

Copied to clipboard

Challenge: Existing research on UI/UX automation often requires high-fidelity inputs like Figma designs or detailed screenshots, limiting accessibility and impeding efficient design iteration.
Approach: They propose a benchmark that evaluates state-of-the-art Vision Language Models on converting sketches into webpage prototypes.
Outcome: The benchmark evaluates state-of-the-art Vision Language Models on automating the conversion of rudimentary sketches into webpage prototypes.
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Vision Language Models struggle with visual arithmetic, seemingly simple tasks like object counting or length comparison, which are essential for relevant complex tasks like chart understanding and geometric reasoning.
Approach: They propose a novel post-training strategy inspired by Piaget’s theory of cognitive development that trains VLMs to recognize invariant properties under visual transformations.
Outcome: The proposed approach outperforms supervised fine-tuning methods while requiring 60% less training data.
Multimodal Fact-Checking with Vision Language Models: A Probing Classifier based Solution with Embedding Strategies (2025.coling-main)

Copied to clipboard

Challenge: Existing fact-checking systems that use text and image information are susceptible to fake news spread by social media platforms.
Approach: They propose a neural probing classifier based on multimodality and embeddings from text and image encoders to represent multimodal content for fact-checking.
Outcome: The proposed classifier outperforms KNN and SVM baselines in leveraging extracted embeddings, highlighting its effectiveness for multimodal fact-checking.
Benchmarking Vision Language Models for Cultural Understanding (2024.emnlp-main)

Copied to clipboard

Challenge: Recent multimodal vision-language models have shown impressive performance in tasks such as image-to-text generation, visual question answering, and image captioning.
Approach: They propose a visual question-answering benchmark to assess VLMs' cultural understanding of various facets of culture from 11 countries across 5 continents.
Outcome: The visual question-answering benchmark aims to assess VLMs' cultural understanding across regions.
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch (2026.acl-long)

Copied to clipboard

Challenge: Existing open-source vision language models lack high-quality training data for chart reasoning . current models are simplistic and repetitive, while associated QA pairs are prone to hallucinations .
Approach: They propose a framework to synthesize complex charts and reliable reasoning data from scratch.
Outcome: Experimental results show that ChartVerse-8B surpasses existing models in QA and difficulty . lack of high-quality training data hampers development of open-source models .
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing (2025.emnlp-main)

Copied to clipboard

Challenge: Current models struggle with long-form videos due to the quadratic complexity of attention mechanisms.
Approach: They propose a model-agnostic framework that leverages temporal cues from queries to prune video tokens.
Outcome: The proposed framework reduces computation by 65% while preserving 97-99% of original performance.
PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of PII leakage ignore how a subject’s online presence affects privacy alignment.
Approach: They propose a benchmark that evaluates safety through the continuum of online presence by stratifying 200 subjects into four visibility categories: high, medium, low, and zero.
Outcome: The proposed model stratifies 200 subjects into four visibility categories based on the extent and nature of their information available online.
Beyond Screenshots: Evaluating VLMs’ Understanding of UI Animations (2026.findings-acl)

Copied to clipboard

Challenge: Recent studies of Vision Language Models (VLMs) for UI understanding have focused primarily on static screenshots, leaving it unclear how well these models handle dynamic UI animations.
Approach: They evaluate UI animation models' ability to perceive animation effects and interpret animation meaning . they use motion, context, and perceptual cues to probe factors affecting VLM performance .
Outcome: The proposed model can detect primitive motion, but its interpretation is inconsistent . the proposed model is based on 300 annotated UI animation videos .
Cultivating Gaming Sense for Yourself: Making VLMs Gaming Experts (2025.acl-long)

Copied to clipboard

Challenge: Recent efforts leverage Vision Language Models (VLMs) as direct controllers, often pausing the game to analyze screens and plan action through language reasoning.
Approach: They propose a paradigm shift in gameplay agent design that uses Vision Language Models as a developer instead of direct control.
Outcome: The proposed framework achieves fluent gameplay in diverse genres, including ACT, FPS, and Flappy Bird, setting a new benchmark for game-playing agents.
Grounding Task Assistance with Multimodal Cues from a Single Demonstration (2025.findings-acl)

Copied to clipboard

Challenge: RGB video often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences embedded in human behavior.
Approach: They propose a framework that integrates eye gaze and speech cues to improve conversational agents for task assistance by integrating eye gaze with speech cuests.
Outcome: The proposed framework captures fine-grained intent and user-specific cues, enabling richer contextual grounding for visual question answering.
VADE: Visual Attention Guided Hallucination Detection and Elimination (2025.findings-acl)

Copied to clipboard

Challenge: Vision Language Models (VLMs) are prone to hallucinations, generating outputs that lack grounding in the actual visual data.
Approach: They propose a sequence modelling approach to learn complex sequential patterns from transformer attention maps.
Outcome: The proposed approach achieves an average PR-AUC of 80% in hallucination detection on M-HalDetect and an 5% improvement in hallucinosis mitigation on MSCOCO.
Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing work on 3D radiograph report generation focuses on 2D images, but 3D medical images provide more comprehensive diagnostic information.
Approach: They propose a comprehensive training recipe for building high-performing VLMs for 3DRRG using a publicly available 3D CT-report dataset.
Outcome: The proposed model achieves superior performance across different model sizes and input 3D medical image resolutions.
Iterative Prompt Refinement for Safer Text-to-Image Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing safety methods for text-to-image models ignore the images produced . this can result in unsafe outputs or unnecessary changes to already safe prompts .
Approach: They propose an iterative prompt refinement algorithm that uses Vision Language Models to analyze prompts and generated images.
Outcome: The proposed method improves safety while maintaining user intent and reliability comparable to existing methods.
Granular Privacy Control for Geolocation with Vision Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Vision Language Models (VLMs) are rapidly advancing in their capability to answer information-seeking questions.
Approach: They develop a benchmark to evaluate the ability of VLMs to moderate geolocation dialogues with users.
Outcome: a new benchmark evaluates the ability of VLMs to moderate geolocation conversations with users.
GUICourse: From General Vision Language Model to Versatile GUI Agent (2025.acl-long)

Copied to clipboard

Challenge: Graphical User Interfaces (GUIs) are a pivotal medium for human-computer interaction.
Approach: They propose a series of datasets for training visual-based GUI agents using general VLMs.
Outcome: The proposed GUICourse datasets show that even a small-sized GUI agent performs better on GUI tasks.
VIBE: Can a VLM Read the Room? (2025.findings-emnlp)

Copied to clipboard

Challenge: Vision Language Models (LLMs) cannot account for the role that non-verbal cues play in understanding social situations.
Approach: They propose a task to test the capabilities of Vision Language Models (VLMs) to account for the visual social-pragmatic inference gap.
Outcome: The proposed task tests the capabilities of a VLM for a social reasoning task.
RG-VQA: Leveraging Retriever-Generator Pipelines for Knowledge Intensive Visual Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve the reasoning capabilities of VQA systems are limited due to complexity of graph neural networks and end-to-end training.
Approach: They propose a method to integrate Dense Passage Retrievers with Vision Language Models to boost the reasoning capabilities of VQA systems.
Outcome: The proposed method outperforms human accuracy and GPT-4 in the ScienceQA dataset.
Unlocking Human-Like Visible Logic: How Logic Diagrams Boost Logic Reasoning in Large Language Models? (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated their remarkable capabilities in natural language understanding and generation, but they struggle with formal logical reasoning.
Approach: They propose to incorporate visual logic diagrams into LLMs’ reasoning workflows to enhance their performance on formal logic tasks.
Outcome: The proposed model improves on syllogistic and conditional reasoning with programmatically generated Venn, Euler, and Linear diagrams.
VISaGE: Understanding Visual Generics and Exceptions (2025.emnlp-main)

Copied to clipboard

Challenge: atypical evaluation instances disrupt incontext instance understanding and in-weight conceptual knowledge.
Approach: They propose to use a dataset to analyze atypical visual and textual images to test their models.
Outcome: The proposed model is based on a dataset consisting of typical and exceptional images.
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches for grounding events in videos are limited by their time-sensitive nature . arrow of time in physics characterizes intrinsic directionality of temporal processes .
Approach: They propose a framework that explicitly models temporal directionality in events to improve event grounding and temporal understanding in VLMs.
Outcome: The proposed framework improves event grounding and directionality understanding in VLMs.
Bears, all bears, and some bears. Language Constraints on Language Models’ Inductive Inferences (2026.findings-acl)

Copied to clipboard

Challenge: Language places subtle constraints on how we make inductive inferences.
Approach: They propose to use language to constrain inductive inferences by replicating an experiment . they find subtle differences arise in general purpose statistical learners like VLMs .
Outcome: The proposed model can be used to extend inductive inferences to humans using language . the model can extend properties of a category to other members of the population, the authors show .
GeoRC: A Benchmark for Geolocation Reasoning Chains (2026.acl-long)

Copied to clipboard

Challenge: Vision Language Models (VLMs) are good at recognizing the global location of a photograph but are startlingly bad at explaining which image evidence led to their location prediction.
Approach: They propose a benchmark for geolocation reasoning chains based on the global location prediction task in the popular GeoGuessr game.
Outcome: The proposed benchmark compares LLM-as-a-judge and VLM-As-jumble strategies against human scoring.
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Recent work on Chain-of-Thought prompting imposes substantial computational overhead . lack of supervision obscures the analyzability of the latent reasoning chain.
Approach: They propose a framework to render latent reasoning chain into images, making latent rationale explicit and traceable.
Outcome: The proposed framework achieves 3-4 token compression and substantial inference acceleration compared to explicit CoT prompting.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations